Biological Imaging
◐ Cambridge University Press (CUP)
Preprints posted in the last 30 days, ranked by how well they match Biological Imaging's content profile, based on 15 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.
Alirezazadeh, P.; Kirsch, E. M.; Tian, Y.; Bewersdorf, J.; Rittscher, J.; Mergenthaler, P.
Show abstract
Speckle artifacts and isolated foreground pixels are common in fluorescence microscopy and can interfere with segmentation and subsequent quantitative image analysis. Conventional denoising methods often modify image intensities through filtering or smoothing, potentially altering biologically relevant fluorescence signals. We introduce Sparse Pixel Cluster Cleaning (SPC-Clean), a topology-aware method that removes poorly supported foreground pixels through iterative neighborhood analysis of a thresholded mask. SPC-Clean is deterministic, training-free, preserves original fluorescence intensities for practical microscopy workflows.
Bhattiprolu, S.; Toor, M.; Soyer, S.
Show abstract
Modern biological imaging generates large, complex datasets that require scalable and reproducible image analysis methods. Deep learning has demonstrated strong performance on bioimage segmentation tasks, but training custom models has remained inaccessible to many researchers due to requirements for GPU infrastructure, programming expertise, and large annotated training datasets. ZEISS arivis Cloud is a browser-based platform for deep learning model training that addresses these barriers through partial annotation support, AI-assisted labeling with SAM (Segment Anything Model), pretrained model initialization, and automatically configured training pipelines requiring no machine learning expertise. The platform supports two segmentation tasks: semantic segmentation using a U-Net-style architecture with an EfficientNet encoder and PixelShuffle decoder, and instance segmentation based on Mask2Former with a Swin-Tiny backbone. Both pipelines incorporate microscopy-specific adaptations including smooth tiling, multi-channel input support, dataset-specific normalization, and partial-annotation-aware loss functions protected by patents US-20240078681-A1 and US-20250111519-A1. Trained models integrate directly with ZEISS arivis Pro for pipeline-based image analysis, ZEISS arivis Hub for parallel execution across large datasets, and ZEISS ZEN for content-aware guided acquisition. We describe the platform architecture, training methodology, segmentation architectures, reproducibility and versioning mechanisms, and FAIR compliance, and illustrate the complete workflow through two intestinal organoid imaging examples. arivis Cloud is freely accessible to student users; other users access the platform via subscription at https://www.arivis.cloud/.
Venturelli, L.; Jacobs, J.; Sifrim, A.
Show abstract
SummaryIntegrating spatial multi-omics data requires coordinated preprocessing, cross-modality alignment and feature registration across modalities that differ in file format, coordinate system and spatial resolution. No existing tool addresses this pipeline end-to-end from raw experimental files till aligned data object. We present FOCUS, an open-source Python package that takes raw data from spatial transcriptomics, mass spectrometry imaging, Raman spectroscopy imaging and brightfield or fluorescence microscopy through modality-specific preprocessing, interactive spatial alignment and resolution-matching registration to a unified MuData object, driven by a single configuration file. Its modular, registry-based architecture allows straightforward extension to additional modalities. FOCUS is accessible via a command-line interface, a browser-based GUI and a Python API. Availability and implementationFOCUS is implemented in Python 3.11, with a browser-based GUI built on a Vue.js 3 frontend served by a Flask backend. Source code, documentation and container recipes are available at https://github.com/sifrimlab/FOCUS; a versioned release is archived on Zenodo (10.5281/zenodo.21700038). Outputs use the AnnData and MuData formats and are directly compatible with the scverse ecosystem.
Wu, Y.-L.; Liu, C.-S.; Lenoir, B.; Merz, K.; Dill, M. T.; Hartmann, F. J.
Show abstract
Summary Multiplexed imaging and spatial proteomics generate complex datasets that require both computational analysis and visual inspection. However, these tasks mostly occur in separate environments because interactive viewers generally require a local display or an additional data server beyond the remote Jupyter sessions itself where large datasets are computationally analyzed. We here present UELer, an interactive viewer that links multi-channel image views with quantitative analysis results directly within Jupyter notebooks, requiring no dedicated infrastructure beyond the notebook session. Cells selected through computational analysis and summary plots can be inspected directly in their tissue context, and selections made in the image can be made available to any downstream analysis. Together, these capabilities support interactive data exploration, iterative cell annotation, and reproducible retrieval of selected regions. Availability and Implementation UELer is a Python package built on ipywidgets and runs in Jupyter environments supporting ipywidgets 8.1 or later, tested in JupyterLab and Visual Studio Code on Linux, macOS, and Windows. It is freely available under GPL-3.0 license and can be installed via pip. Source code and documentation are available at https://github.com/HartmannLab/UELer and https://hartmannlab.github.io/UELer/. An online, no-install version runs remotely via BinderHub (https://mybinder.org/v2/gh/HartmannLab/UELer/main), accessible through the script/run_ueler_binder.ipynb notebook.
Fernandez, N.; Ishar, J.; Wang, H.; Saad, A. B.; Lipinski, M.; Farhi, S. L.
Show abstract
Spatial-transcriptomics integrates high-dimensional single-cell data with microscopy to reveal cellular states, communication, and tissue organization. Analyzing this data requires a combination of multi-modal data processing, high-dimensional data analysis, spatial analysis, and integrated visualization. However, computational analysis is increasingly becoming a bottleneck as approaches mature and dataset sizes increase. Additionally, visualization can be challenging as open-source visualization tools struggle to scale to large datasets (exceeding 1 billion transcripts), and commercial visualization tools are costly, closed source, and inflexible. We present Celldega, an open-source Python and JavaScript library for scalable, interactive visualization and analysis of spatial-omics data. Celldega integrates custom analyses, performs neighborhood analysis, implements an efficient visualization-specific file format, and enables interactive exploration in notebooks and web galleries. We demonstrate Celldega across multiple technologies, tissues, and datasets, including 3D reconstructions of the developing whole mouse head comprising over four million cells. Finally, we demonstrate how Celldega can be utilized throughout the entire lifecycle of spatial data analysis, from quality control to building a public shareable gallery.
Park, J.; Ratka, M.; Biswas, A.; Shofner, I.; Kerns, K.; Sarkar, A.
Show abstract
Reliable delineation of the head and tail of swine spermatozoa supports automated assessment of boar semen quality, from morphometric measurement to the quality control of insemination doses. In practice this relies on fluorescent staining, which adds chemistry, cost, and delay to every acquisition and labels only the nucleus. Recent work coupling imaging flow cytometry with machine learning has advanced rapidly, yet the segmentation stage still depends on a stained channel at inference and resolves the head alone. We present a supervised encoder decoder network that segments boar spermatozoa from brightfield images acquired on an Amnis ImageStream Mark II with no stain at inference. Training labels derive from the Hoechst 33342 nuclear channel (Ch7), recorded in registration with brightfield (Ch1); the dye serves only as an annotation source, and the network sees Ch1 alone. The best semantic segmentation model reaches a Dice coefficient of 0.940 on held-out cells. For comparison we evaluate a classical morphological pipeline, four further semantic segmentation models spanning three decoder families and two ImageNet-pretrained backbones, and two zero-shot pipelines built on the Segment Anything Model 2 (SAM 2), prompted either by a dilated box around the predicted head mask or by head and tail boxes emitted by a Gemma 4 Vision Language Model (VLM). The zero-shot route scores 0.637 against Ch7 but labels the tail, which the fluorescence protocol cannot. Cells scoring worst under the supervised model proved to be mostly registration failures rather than segmentation failures, as Ch7 is displaced relative to Ch1. Manual screening for this drift is infeasible at dataset scale, so we propose a flagging system that marks any Dice below 0.792, two standard deviations below the mean, and pairs it with a zero-shot pipeline in which a VLM l and SAM 2 cross-check the flagged cell before human review.
Kesenci, Y.; Le Folgoc, L.; Angelini, E.
Show abstract
Deep-learning-based segmentation algorithms have gained considerable accuracy for processing biological images. In particular, the introduction of large foundation models, novel architectures, and semantically varied datasets now allows for deployment of state-of-the-art models for clean image cohorts with limited re-training or, in the best of cases, in an out-of-the-box fashion. Biological imaging, however, is liable to corruptions that can hinder their deployment. While some methods document their robustness to the most common corruptions, a systematic robustness analysis of the state of the art to the expansive gamut of corruptions in biological imaging remains to be done. We perform this benchmarking by simulating 36 corruption types with varying degradation severity on images sampled from 30 different datasets. Our benchmark accounts both for the variety in biological images and the nature of corruptions. Among other things, our study reveals that performance on clean images does not correlate with overall robustness to image corruptions. In fact, we find that a decade-old method, StarDist, is more robust than many of its more recent foundation-model-based counterparts. We also show in a dedicated representation analysis that the performance of segmentation models collapses in the early layers of the encoding phase.
Guedes, J.; Sliwa-Gonzalez, A.; Szadai, L.; Geiger, P.; Woldmar, N.; Reyes, M. A.; Bastida, R. A.; Coto, D. L. F.; Oskolas, H.; Marko-Varga, M.; Schultz, L.; Appelqvist, R.; Wieslander, E.; Malm, J.; Marko-Varga, G.; Gil, J.
Show abstract
Melanoma incidence continues to rise globally, with formalin-fixed paraffin-embedded (FFPE) tissue archives representing an invaluable resource for large-scale retrospective proteomic studies. However, inconsistent deparaffinization remains a critical pre-analytical bottleneck limiting protein yield, reproducibility, and downstream data quality. In this study, we developed and validated a fully automated FFPE deparaffinization workflow using the Fluent(R) 780 liquid handling workstation (Tecan (C)) and evaluated its performance against a conventional manual protocol in a cohort of 54 patients with primary cutaneous melanoma, predominantly at early AJCC 8th edition stage I-II. The automated workflow achieved superior protein identification (6,146 {+/-} 860 vs. 4,941 {+/-} 1,091 proteins; p < 0.0001) with lower technical variability, while maintaining highly comparable global proteomic profiles as confirmed by principal component analysis and hierarchical clustering. A total of 8,305 proteins (96.1%) were identified by both methods, supporting the reproducibility and equivalence of the automated approach. Patients were stratified by the presence (N=21) or absence (N=33) of histological regression in the primary tumor. Proteomic comparison revealed 97 upregulated and 226 downregulated proteins in regressing melanomas, with pathway enrichment analysis demonstrating elevated mitochondrial and translational activity alongside reduced innate immune and complement pathway activation in the regression group. No statistically significant differences in overall, disease-free, or progression-free survival were observed between groups, consistent with the early-stage composition of the cohort. Digital pathology validated tissue morphology preservation across processing conditions. These findings support the integration of automated FFPE processing with proteomic and digital pathology workflows as a scalable platform for precision melanoma research. TOC Figure O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=133 SRC="FIGDIR/small/744404v1_ufig1.gif" ALT="Figure 1"> View larger version (49K): org.highwire.dtl.DTLVardef@1d51629org.highwire.dtl.DTLVardef@a1f126org.highwire.dtl.DTLVardef@1df1b0aorg.highwire.dtl.DTLVardef@686f1c_HPS_FORMAT_FIGEXP M_FIG C_FIG
Shtyrov, A.; Wilson, H.; Murshudov, G. N.
Show abstract
Damage to biological specimens by the electron beam is the fundamental resolution-limiting factor in cryoelectron microscopy (cryo-EM) single particle analysis. There is, however, currently no method to accurately infer fluence-dependent changes to the specimen structure during electron irradiation. We develop a Bayesian framework to fit a sequence of atomic models to a series of cryo-EM reconstructions produced at increasing fluence. In particular, our algorithm is able to infer the ensemble average position and atomic displacement parameter of every atom in the macromolecule as a function of fluence. Application of the algorithm to cryo-EM datasets shows that the molecule expands during imaging and identifies environment-dependent variations in beam-induced damage. We use our results to propose a stochastic process model of this phenomenon. We envisage that our method will lead to a better mechanistic understanding of radiation damage to biological specimens and may contribute to efforts to mitigate its effects.
Bourne, R. M.; Arhatari, B.; Watson, G.; Gureyev, T.; Phipps, A.; Dowland, S.; Kurniawan, N.; Sved, P.
Show abstract
Formalin-fixed prostate tissue samples were imaged by propagation-based synchrotron phase contrast micro computed tomography ({micro}CT) with a 3D spatial resolution of ca. 3 {micro}m. Post-{micro}CT, samples were prepared for histology with sections close to coplanar with the transverse {micro}CT image planes. Haematoxylin and eosin stained sections were examined by an expert prostate histopathologist and compared qualitatively with corresponding {micro}CT-visible microstructure features. There is potential for {micro}CT to provide complimentary information to conventional histology and light microscopy without the need for preparation of stained thin sections. For the imaging conditions and spatial resolution of our study, {micro}CT may provide tissue architectural features similar to those used in Gleason grading, albeit without clear subcellular microstructure detail. At the spatial resolution of our study {micro}CT may provide novel 3D microstructure information for validation of diffusion weighted magnetic resonance imaging (MRI) methods. As an example, we demonstrate a qualitative correlation between {micro}CT-derived stromal fibre orientation and preferential water diffusion direction measured by diffusion tensor MRI microscopy of the same sample.
Potter, L. A.; Trull, A.; Kumar, N.; Drake, O. R.; Nogueira, M.; Peters, J.; Heinsbroek, J. A.; Day, J. J.; Worthey, E. A.; Ianov, L.
Show abstract
Recent advances in spatial transcriptomics have enabled the profiling of increasingly larger numbers of genes while retaining single-cell and subcellular resolution in situ. However, standardized bioinformatics workflows for analyzing these datasets have lagged behind, with existing pipelines focusing primarily on image processing and cell segmentation. To address this gap, we present nf_xpatial, a best-practices Nextflow pipeline for the downstream analysis of 10x Genomics Xenium data. The pipeline performs quality control, filtering, log and cell area normalization, multi-sample integration, and both expression-driven and spatially informed clustering across systematic parameter sweeps, allowing users to evaluate and compare clustering resolutions and spatial modeling parameters within a single reproducible run. Overall, nf_xpatial streamlines the processing of Xenium data from platform outputs to integrated single-cell and spatial clustering datasets, providing a standardized starting point from which biologists can fine-tune parameters and proceed to hypothesis-driven spatial analyses.
Moustafa, S.; Zheng, Y.; Rendeiro, A. F.
Show abstract
Stain normalization reduces color variability in histopathology whole-slide images, but cohort-scale pipelines lack fused multi-image batch transforms for classical methods. We present StainX, a GPU-accelerated batch stain normalization framework built around a two-stage fit/transform interface. It implements histogram matching, Macenko, and Reinhard normalizers through a portable PyTorch backend and an optional CUDA backend that fuses per-pixel operations for batch throughput. On NVIDIA GPUs, the fused CUDA path outperforms the torch CPU backend by 168x, 70x, and 48x for Reinhard, histogram matching, and Macenko respectively, and exceeds the fastest GPU peers by 7-8x (Reinhard) and 2x (Macenko) at comparable accuracy. StainX also provides user-selectable precision modes, a documented Python API, continuous integration testing, and online documentation. Source code available at https://github.com/rendeirolab/stainx, and documentation at https://stainx.readthedocs.io. Implemented in Python. Runs on Linux, macOS, and Windows.
Morris, C. A.; Bastian, W. C.; Cui, Y.; Kurago, Z.; Douglass, E. F.
Show abstract
Single-cell spatial transcriptomics can connect molecular cell states with tissue morphology, but this promise depends on accurate registration to histopathology. In serial sections, however, tissue borders often differ because of sectioning artifacts, staining variability, and field-of-view acquisition, limiting conventional area-based registration. We developed AnchorR, an expert-guided workflow for coarse-grained alignment of hematoxylin and eosin (H&E) images with CosMx Spatial Molecular Imaging data. Bioinformaticians first define and color-code cell types in Seurat, and pathologists then identify corresponding internal landmarks using QuPath overlays. AnchorR combines these paired landmarks to estimate affine transformations, quantify residual error, and support visual quality control and anchor refinement. Using six oral pre-cancerous tissue sections, we identified 60 cross-modal landmarks. Fitting each section independently reduced mean landmark error from 121.5 {micro}m with a single whole-slide transformation to 14.6 {micro}m. Cross-validation further showed that increasing the number of anchors improved robustness, with nine-anchor fits achieving approximately 20 {micro}m error, or about one cell diameter. AnchorR is designed to complement automated computer-vision methods by providing reliable tissue-level alignment when border mismatch makes global registration difficult. By creating a shared workspace for pathologists and bioinformaticians, it operationalizes an expert-in-the-loop approach and makes feature-based multimodal registration accessible without specialized computer-vision expertise or high-performance computing.
Bregy, I.; Mesman, R.; Tassan-Lugrezin, S.; Kooij, T. W. A.; van Niftrik, L.
Show abstract
Researchers using electron microscopy must often balance a trade-off between obtaining high-resolution structural information and preserving sufficient cellular context. At one end of this spectrum, single particle cryo-electron microscopy and cryo-electron tomography provide near-molecular detail but are typically limited to relatively small fields of view. At the other, volume electron microscopy approaches, such as scanning electron microscopy of resin-embedded specimens, capture large cellular volumes but generally at lower resolution. Consequently, linking nanoscale structural information to larger cellular architecture remains a significant challenge. To address this gap, we optimised a transmission electron tomography workflow for resin-embedded malaria parasites that allows us to visualise targeted regions of interest at nanometre-scale resolution while retaining several micrometres of surrounding cellular context. Here, we present our current best-practice pipeline for sample preparation, tomogram acquisition, and reconstruction. In addition, we introduce VolWeaver, a data-processing framework, that integrates high-resolution tomographic datasets into serial section volume reconstructions, enabling the visualisation and interpretation of ultrastructural features within their broader cellular environment.
Neumann, M.; Arras, P.; Kaster, A.-K.; Ott, A.
Show abstract
Multimodal Gaussian process factor analysis provides a flexible framework for dimensionality reduction in temporally or spatially resolved omics data. Existing approaches, however, typically rely on pre-specified Gaussian process kernel families and do not explicitly separate each latent factor into a component capturing gradual, smooth variation and a complementary component capturing fine-scale, non-smooth variation. Here, we present MOFTy, a Bayesian multimodal factor analysis framework based on numerical information field theory (NIFTy) that replaces fixed kernel families with the flexible correlated field model in NIFTy and enables explicit additive component separation within each latent factor with quantified uncertainty. NIFTy has been successfully applied to high-resolution Bayesian imaging in astrophysics and facilitates scalable, curvature-aware variational inference for efficient posterior approximations. We validate MOFTy on simulated data; applications to published multi-omics data demonstrate that MOFTy disentangles latent spatial structures by separating smooth gradients from localized fine-scale heterogeneity in human glioblastoma and recovers cross-modal patterns in a mouse gastrulation dataset.
Collins, J. T.; Wang, Q.; Williams, G. O. S.; Stewart, H.; Wood, H. A. C.; Parry, C.; Toogood, C. M.; Bruce, A. M.; Young, V.; Moore, A. M.; Dorward, D. A.; Marshall, A. D. L.; Pellicoro, A.; Bain, L.; Akram, A. R.; Dhaliwal, K.; Stone, J. M.
Show abstract
Background: Accurate sampling of suspected peripheral lung cancers depends on access to the lesion and confirmation that the biopsy tool is in contact with target tissue. Current bronchoscopic navigation and imaging techniques can guide instruments to a target but do not provide real-time biological confirmation at the point of sampling. Fluorescence lifetime imaging microscopy (FLIM) provides molecular contrast by measuring fluorescence decay - how long photons continue to be emitted from fluorescent molecules. In the Precision Lung clinical study (ISRCTN15093468), the Prothea Imaging System (Generation 1) identified a candidate tumour-associated phenotype of spatially overlapped low fluorescence lifetime and low intensity (LLLI) from in-vivo imaging. We used this observation as the basis for a reverse-translational study to determine whether the LLLI phenotype is linked to cancer pathology; reproducible with the Imaging System (Generation 2); and distinguishable from normal lung tissue. Methods: Previously reported Precision Lung findings were used as the clinical starting observation and were not re-analysed. Validation was then performed using: (i) pathology linked benchtop FLIM of early-stage non-small-cell lung tissue microarrays encompassing malignant cell clusters of approximately 300 um2, matched to the EoT imaging scale; (ii) five sequential fresh lung-cancer resections imaged at tumour and comparator regions, including visibly blood-rich contact sites, using the (Generation 2) Imaging System; and (iii) systematic mapping of two ventilated non-cancer donor lungs, one from a smoker and one from a non-smoker, across all available lobes. The LLLI phenotype was defined as spatial co-localisation of low intensity and short lifetime. Results: Using a real time fibre based FLIM system, capable of deployment through a working channel of a bronchoscope, the LLLI tumour phenotype was optically identified in freshly resected tumour tissue. The same phenotype was identified in fixed tissue samples with known pathology, and with images taken in the Precision Lung clinical study. Whole human lung controls did not show evidence of the tumour phenotype. Conclusions: This evidence forms a reverse-translational chain that supports the concept of the Prothea Imaging System - as a platform that confirms that the tool is in contact with a region of cancer in the lesion, while preserving continuous access for biopsy or intervention.
Velazquez, D.; Hallinan, C.; An, R.; Clifton, K.; Fan, J.
Show abstract
Abstract Imaging-based spatially resolved transcriptomics (imSRT) technologies provide high-throughput molecular-resolution spatial characterization of genes within cells. Conventional analysis methods to identify cell-types and states in imSRT data rely on gene count matrices derived from tallying the number of mRNA molecules detected for each gene per segmented cell, thereby overlooking subcellular heterogeneity that can be useful in defining cell states. To take advantage of the molecular-resolution information in imSRT data and potentially identify cell-states based on subcellular heterogeneity, we developed STARIT (Spatial Transcriptomics As Rasterized Image Tensors). STARIT converts transcripts within segmented cells in imSRT data into an image-based tensor representation that can be combined with deep learning computer vision models for downstream analysis. Using simulated and real imSRT data, we demonstrate that STARIT distinguishes transcriptionally distinct cell-types and further separates cell states based on subcellular transcript localization, which conventional gene count analysis fails to capture. By providing a standardized framework to encode subcellular molecular information in imSRT data, STARIT will enable deeper insights into subcellular heterogeneity and enhance the identification and characterization of cell-types and states that are overlooked by gene count representations.
Poirier, C.; Petit, L.; Lefebvre, J.; Descoteaux, M.
Show abstract
To disentangle complex fiber configurations that remain challenging for diffusion MRI tractography, insights might be gained from microscopy tractography. Indeed, by precisely following small white matter (WM) fascicles invisible at the resolution of diffusion MRI, microscopy tractography can help explain how fiber populations are organized at the finest scales. Serial optical coherence tomography (S-OCT) is an imaging modality relying on the intrinsic contrast of a sample. When applied to brain tissues, the S-OCT contrast is primarily driven by the myelin reflectivity. Due to its high resolution, on the order of microns, and its 3D nature, S-OCT offers promise for studying WM connections at the microscale. However, while other microscopy imaging modalities have been shown to enable tractography, whether the reflectivity contrast from S-OCT supports the reconstruction of long-range WM fascicles at the microscale remains unknown. Furthermore, there is a gap in the literature regarding how an ideal microscopy tractography algorithm should behave with respect to the choice of tractography algorithm, tracking maps definition and microscale orientation distribution functions (ODF) estimation. In this work, we describe a tailored approach to reconstruct WM fascicles at the microscale from S-OCT acquisitions. We improve microscale orientation distribution functions (ODF) estimation by implementing a sliding-window formulation allowing the estimation of ODF at S-OCT resolution, and use apodized Dirac delta functions for reducing unwanted interference. We validate our approach on a simulated microscopy-like FiberCup dataset, and show that using multiscale Frangi filters for estimating ODF outperforms structure tensor analysis. We also show that particle filtering tractography with anatomical constraints enables targetted, region-to-region tractography, and outperforms standard deterministic or probabilistic tracking approaches. We further demonstrate our method on a whole mouse brain S-OCT reconstruction at 10 m by reconstructing the thalamocortical white-matter projections. Overall, our results show that S-OCT tractography recovers fine white matter fascicles visible at the microscale, and that these connections are supported by viral tracing experiments from the Allen Mouse Brain Connectivity Atlas. Moreover, this work shows the first ODF estimation and fully-3D probabilistic particle filtering tractography of the mouse brain from S-OCT reconstructions at 10 m isotropic resolution.
Kharbanda, M.; Tubelleza, R.; Tan, Y.; Tan, C. W.; Janke, C.; Sebina, I.; Belz, G.; Kulasinghe, A.; Salim, A.; Bhuva, D. D.
Show abstract
Multiplex spatial proteomics enables highly multiplexed in situ profiling but remains limited by technical variation arising from autofluorescence, non-specific antibody binding, instrument noise, and staining variability, affecting downstream biological tasks like cell typing. We present bgnorm, a statistical framework that describes fluorescence measurements using a generative mixture model of background, non-specific binding, and biological signal components. As natural statistical consequences, the model yielded three new methods: a background-correction method through probabilistic deconvolution of protein intensities, quality control metrics, and a quantile normalisation approach to unify measurements across markers, samples, and sequential slices. Across multiple multiplex imaging technologies, bgnorm improves signal separation and downstream marker positivity classification compared with existing preprocessing approaches. In expert-annotated datasets comprising over 406,000 marker positivity annotations, bgnorm achieved the highest classification performance and enabled accurate use of a single global positivity threshold across markers and samples. The method is implemented in the bgnormR and bgnormpy packages.
Carlsen, A. S.; Chen, T.; Cowie, N. L.; Brinch, C.; Groves, T.; Nielsen, L. K.
Show abstract
Isotopic Metabolic Flux Analysis (I-MFA) is a standard approach for estimating intracellular metabolic fluxes. I-MFA infers fluxes by comparing simulated and measured metabolite isotopologue distributions (MIDs) of metabolites from isotope labeling experiments. MIDs represent fractional abundances that strictly sum to one for any given metabolite, thus they are inherently compositional data. However, state-of-the-art estimation approaches rely on calculating standard Euclidean distances between MIDs in a non-compositional paradigm, introducing a systemic bias. To resolve this, our study proposes compositional I-MFA. We demonstrate how to construct a meaningful orthonormal basis for MIDs via ordered sequential binary partitioning, which can be used to perform isometric log-ratio (ILR) transformation. As a minimal change to existing I-MFA workflows, we suggest estimating fluxes by minimizing Euclidean distances between ILR-transformed MIDs. We validated this framework against traditional methods using both a toy model and a biologically realistic model, evaluating point estimates, sensitivity across varied true fluxes, and confidence intervals. In the two examples, compositional I-MFA consistently outperformed traditional approaches, reducing mean squared error of flux point estimates by an average of 42.6% and substantially narrowing confidence intervals. We conclude that compositional data analysis significantly improves I-MFA and can be implemented as a simple drop-in replacement for current pipelines. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=156 SRC="FIGDIR/small/742769v1_ufig1.gif" ALT="Figure 1"> View larger version (34K): org.highwire.dtl.DTLVardef@1a53aa4org.highwire.dtl.DTLVardef@ad225aorg.highwire.dtl.DTLVardef@aa430eorg.highwire.dtl.DTLVardef@1880ca_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LINew compositional data approach improves metabolic flux estimation. C_LIO_LIThis data transformation requires minimal changes to existing workflows. C_LIO_LIThe new method reduced MSE of flux estimates by 42.6% in two examples tested. C_LIO_LIThe confidence intervals of the estimated fluxes were substantially narrowed. C_LIO_LIEstimation accuracy remained robust across a wide range of metabolic fluxes. C_LI